Papers with text-to-SQL generation

10 papers
VeriMinder: Mitigating Analytical Vulnerabilities in NL2SQL (2025.acl-demo)

Copied to clipboard

Challenge: Application systems using natural language interfaces to databases (NLIDBs) have democratized data analysis, but they are not without significant risks.
Approach: They propose an interactive system that detects and mitigates cognitive biases in analytical questions by using contextual semantic mapping frameworks.
Outcome: The proposed system detects and mitigates cognitive biases in analytical questions and generates high-quality, task-specific prompts.
PExA: Parallel Exploration Agent for Complex Text-to-SQL (2026.acl-short)

Copied to clipboard

Challenge: Recent work in text-to-SQL has explored toolaugmented LLMs, deep planning, and agentic workflows to address complex challenges.
Approach: They validated a framework for text-to-SQL, Spider 2.0, with 70.2% execution accuracy.
Outcome: The proposed framework achieves 70.2% execution accuracy on a state-of-the-art benchmark for text-to-SQL, Spider 2.0.
SeaD: End-to-end Text-to-SQL Generation with Schema-aware Denoising (2022.findings-naacl)

Copied to clipboard

Challenge: Using sketch-based slot filling, text-to-SQL models suffer from over-complexity . et al., e.al., and d.albert, dr., propose a novel method for text- to-Sql generation .
Approach: They propose to train sequence-to-sequence model with Schema-aware Denoising . they propose a clause-sensitive execution guided (EG) decoding strategy .
Outcome: The proposed method improves performance in schema linking and grammar correctness . it also establishes new state-of-the-art on the WikiSQL benchmark .
TabSQLify: Enhancing Reasoning Capabilities of LLMs Through Table Decomposition (2024.naacl-long)

Copied to clipboard

Challenge: Large language models struggle with large tables due to their limited input length . a novel method that decomposes tables into smaller and relevant sub-tables reduces the computational load on LLMs .
Approach: They propose a method that leverages text-to-SQL generation to decompose tables into smaller and relevant sub-tables . the method can reduce the input context length significantly, making it more scalable and efficient .
Outcome: The proposed method performs remarkably well on the WikiTQ benchmark and on the TabFact benchmark.
Data-Anonymous Encoding for Text-to-SQL Generation (D19-1)

Copied to clipboard

Challenge: Existing approaches to handle table-related tokens before the semantic parser are not efficient . existing approaches ignore handling table- related tokens or use deterministic approaches based on string-match or word embedding similarity.
Approach: They propose a more efficient approach to handle table-related tokens before the parser . they propose tagging a sequential tabbing problem and an implicit supervision approach .
Outcome: The proposed approach significantly outperforms deterministic approaches.
Multitask Pretraining with Structured Knowledge for Text-to-SQL Generation (2023.acl-long)

Copied to clipboard

Challenge: Existing methods for learning representations of structured knowledge are limited to the minority of people with technical skills.
Approach: They propose a large pretraining dataset and strategy for learning representations of text, tables, and SQL code that leverages the entire context of the problem.
Outcome: The proposed model improves on two SQL tasks and shows a 1.7 and 2.2 percentage point improvement over existing methods.
Clause-Wise and Recursive Decoding for Complex and Cross-Domain Text-to-SQL Generation (D19-1)

Copied to clipboard

Challenge: Existing deep learning approaches for text-to-SQL generation are limited to the WikiSQl dataset . a novel clause-wise decoding neural network model can be used to generate complex queries over multiple databases .
Approach: They propose a SQL clause-wise decoding neural architecture with a schema encoder to address the Spider task.
Outcome: The proposed model achieves 4.6% accuracy gain on the Spider dataset and 9.8% accuracy gain in test and dev sets.
STaR-SQL: Self-Taught Reasoner for Text-to-SQL (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for generating step-by-step “chain-of-thought” rationales are limited to text-to-SQL.
Approach: They propose a method that prompts SQL query generation to produce reasoning steps for SQL queries and fine-tunes it on rationales that lead to correct outcomes.
Outcome: The proposed method outperforms agent-like prompting methods on the Spider benchmark.
TinySQL: A Progressive Text-to-SQL Dataset for Mechanistic Interpretability Research (2025.emnlp-main)

Copied to clipboard

Challenge: Existing text-to-SQL datasets are too complex and noisy for rigorous interpretability analysis.
Approach: They propose text-to-SQL generation as an ideal task to study mechanistic interpretability . they use edge attribution patching and sparse autoencoders to identify minimal circuits .
Outcome: The proposed task combines the formal structure of toy tasks with real-world complexity.
VET: Verifiable Execution Tracing for Reliable Text-to-SQL Generation (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for text-to-SQL generation are prone to hallucinations and grounding . authors present a novel reasoning paradigm that transforms text- to-Sql from unverifiable textual rationales into step-wise executable semantics.
Approach: They propose a reasoning paradigm that transforms text-to-SQL from unverifiable textual rationales into step-wise executable semantics.
Outcome: The proposed reasoning paradigm transforms text-to-SQL from unverifiable textual rationales into step-wise executable semantics.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations